social science eur
The advantages of context specific language models: the case of the Erasmian Language Model
Gonçalves, João, Jelicic, Nick, Murgia, Michele, Stamhuis, Evert
The dominant The current trend to improve language trend in improving language models model performance seems to be based on seems to be based on scaling up the scaling up with the number of parameters number of parameters (e.g. GPT4 has (e.g. the state of the art GPT4 model has approximately 1.7 trillion parameters) or approximately 1.7 trillion parameters) or the amount of training data fed into the the amount of training data fed into the model (e.g. the Common Crawl dataset model. However this comes at significant currently has more than 75 TB of textual costs in terms of computational resources data). However, this scaling usually comes and energy costs that compromise the sustainability at significant costs in terms of computational of AI solutions, as well as risk relating resources, which translate to energy to privacy and misuse. In this paper (Emma, Ananya, & Andrew, 2019), financial, we present the Erasmian Language Model and environmental costs, and data collection, (ELM) a small context specific, 900 million which in turn raise concerns regarding parameter model, pre-trained and finetuned legal, privacy, quality, and responsibility by and for Erasmus University Rotterdam.